STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress - #1013
Draft
MauroToscano wants to merge 1323 commits into
Draft
MauroToscano wants to merge 1323 commits into
MauroToscano wants to merge 1323 commits into
Conversation
recompute_lde_produces_byte_identical_proofs compares the bytes of two proves of one instance, one per residency mode, at the test options' grinding factor of 1. Under `parallel` the CPU nonce search is rayon's find_any (crypto::grinding::generate_nonce), so the two proves can return different valid nonces; the nonce is absorbed before the queries are drawn, and every opening after it moves. The test fails whenever the two searches disagree, whichever residency mode runs: on a laptop it failed 11/20 at d1dc455 and 15/20 at 7d41668, and 20/20 passed with LAMBDA_VM_DETERMINISTIC_GRIND=1 (the smallest nonce) or without `parallel` (a sequential find). The residency tests now prove at grinding factor 0, as zf_golden_tests already does for the same reason. The residency mode acts on the main LDE, which the grind never reads, so the comparison loses nothing it could catch.
The residency-mode tests compared two proves byte for byte while grinding with rayon's find_any, so any two nonce searches could disagree. They now prove at grinding factor 0, as the golden tests do.
Every STARK wrap folded its program_id in-guest with one keccak permutation. That permutation keeps the whole keccak family (LFM_KECCAK, KECCAK_RND, KECCAK_RC) and, through its byte lookups, BITWISE in every wrap, and makes every level-1 node re-verify those four sub-proofs per child. By default the wrap now computes the id at emission with recursion::program_id_from_digest (still keccak, the SOUNDNESS.md 6.7 carve-out) over the values derived from the trusted ELF, publishes it as program text in the fold's two-word layout, and binds every input the fold consumed with an equality assert on the cell the verification reads: the ELF digest halves the statement absorbs, the DECODE root Phase A absorbs and the DECODE leg compares, pc_start, and the page roots. A proof over any other value has no execution. The constants are LFM_CONST rows, so they are in the wrap's program_id, which its parent interns: the wraps, and every node and root above them, become functions of the ELF, as the WHIR wraps already are. SOUNDNESS.md 6.9 states what binds each input. LAMBDA_VM_STARK_WRAP_FOLD=1 keeps the in-guest fold. Its branch is the previous emission unchanged, so every wrap, node and root program is today's byte for byte. The setting is read once per process and named on stderr; tests override it per thread. With no keccak left the wrap's mask drops the keccak family and, under RPX, BITWISE: no chip it instantiates sends BITWISE a lookup. On block 25368371 the census falls by 1.05 G cells (15 wraps -26.3 M each, level 1 -783 M, levels 2-4 +127 M). The WHIR programs do not read the setting. Tests. Laptop: the standalone attestation at both root widths publishes the fold's words, refuses a forged constant for each field and every tampered cell, and leaves no BITWISE sender without its receiver. Box: on the real epoch the default wrap publishes the fold's words, emits no keccak, carries the masks above, and refuses a forged ELF digest, pc_start or DECODE constant and a tampered cell; the WHIR wrap's program is the same under both settings.
…evice halves The artifact build (`build_artifacts_with_hasher`) now walks its commits through `lfm::artifact_walk`: one plan (the eleven slot groups, the LFM_BLAKE3 chunks, the LFM_HASH tail, under row-pair and, when the format has it, one-row leaves) and one walk parameterized by a pass. `Pass::All` is the build exactly as it ran: the same windows of `groups_in_flight`, the same just-in-time BLAKE3 chunk materialization, the same `commit_group_device_or_host_with` per group. New, and unused by any caller yet: `build_artifacts_with_device_section` (and `build_artifacts_sectioned`, which also returns the split). It walks `Pass::Host` first — the groups the device would decline, committed by `commit_group_host_with`, which cannot reach the card — and then enters a caller-supplied section (a card permit) for `Pass::Device`, in the same windows minus the host groups. The merge refuses a slot committed on both sides or on neither. Routing is `gpu_lde::commit_reaches_device`, the admission `admit_commit` applies (no device or below the row floor declines; over budget still reaches the device and aborts there). The registry drift tests pin the default walk's roots; the merge's two refusals are unit-tested.
…t field A WHIR round has three proof-of-work slots (folding, out-of-domain, query), and the proof carries three nonces a round whatever the bits. A slot whose grind has zero bits, and the last round's out-of-domain slot, is carried and never read, so any value in it verifies. A query-only grind (P2) would leave two such unbound fields a round. ChainFormat gains `nonces: NonceLayout`: - Three (the default): today's format, byte for byte. - Spent: a round carries only the nonces its grinds spend. The in-guest arena has no word for an unspent nonce; ChainShape::carries is the one place the layout is written, and the word count, the hints and the arena words all read it. The host verifier refuses a nonzero value in a host field the layout does not carry (Error::UnspentNonce). RoundNonces keeps its three fields, so both layouts share one proof type and Three keeps its bytes. GrindBits::query_only(bits) grinds before the query positions only. The query count reads the query grind alone, so it does not move. No production config uses Spent or a query-only grind yet, so every proof, program and pin is unchanged. The legacy layout's chain programs, arenas and proof bytes are pinned against values printed at 0428c39, and the transcript closed form now prices only the grinds a config spends.
…t off) STARK level 0 idles the card 12.9 s of 79.3 s (G1): 5.44 s inside build_artifacts holds (64 % idle), 4.95 s inside multi_prove holds (30 %), 2.85 s with no hold. The interior's holds, on larger programs, are 10 % and 13 % idle; what differs at level 0 is the host load of the other five workers (the epoch reconstructs above all). Three knobs, in `lfm::card_schedule`, each moving work and never a committed byte: - LAMBDA_VM_GAP_PREP_SCOPE=1: `build_artifacts_counted` holds the card only around the build's device commits (`build_artifacts_sectioned`); the groups under the device floor are committed on the host first. With the card trace on, each build prints its host/device split. - LAMBDA_VM_GAP_PREP_NICE=<1..19>: host-only phases run on a rayon pool of their own whose threads take that nice value (Linux setpriority; libc as a Linux-only dependency): `lfm_prepare`'s execute and fill, the sectioned build's host half, and the STARK tree's wrap reconstruct, emit, census and harvest and the global child's harvest, emits and host verifies. The card holder keeps the CPU and the global pool. - LAMBDA_VM_GAP_PREP_AHEAD=1: a level-0 wrap builds its artifacts on a helper thread while it executes and fills (`lfm_prove` is now `lfm_prepare` + `lfm_prove_prepared`; nothing before multi_prove reads the artifacts). Costs host memory: traces exist while the build may queue. The permit stays a mutual exclusion under all three, and a build's device commits run in the windows they always did. Unset, every path is the old one; the STARK driver prints a CARD SCHEDULE line only when a knob is set. Tests: the sectioned build equals the whole build over four programs (split hash, three BLAKE3 chunks) and three one-row modes, and enters its section once exactly when it has device work; on cuda, every device commit lands inside the section. Execute and fill on the pool are cell-identical to inline over the trace-identity cases; the pool's threads carry the nice value (Linux). The prepared prove publishes the same words and verifies (box-scale), refuses another hasher's artifacts, and the driver's AHEAD helper returns the plain build's artifacts and a verifying proof (box-scale).
Each WHIR base-chain round ground 20 bits before three challenges. Only the query grind buys proven bits as placed: the folding grind sits before the round's first sumcheck message, so the first folding challenge is redrawn by varying that message at one hash a try, and the out-of-domain grind follows the out-of-domain point. The new ZF lever `whir_grind` therefore defaults to `query`: GrindBits::query_only(20) under NonceLayout::Spent, one grind and one nonce word a round, 518 grinds a block instead of 1,472 at stack 27. The query count reads the query grind alone and stays 112. The proven bits per phase do not move: chain minimum 130.393 at stack 27, pipeline minimum 128.946. LAMBDA_VM_ZF_WHIR_GRIND=all is the opt-out: GrindBits::uniform(20) under NonceLayout::Three, the production config from before this commit. A test pins it against a literal, and its chain programs against the values printed at 0428c39. The banner gains `whir_grind=`. No univariate option reads the lever, so no STARK proof, program or id moves. Re-blessed: the production default chain's pins now describe the P2 chain (12 grind permutations, 16,411 permutations, 32,590 arena words, 150,258 / 202,873 rows). Its previous pins move unchanged to the opt-out's test. The banner strings in zf_format's tests gain the new key.
…head of it The first prove of a process builds every domain and twiddle set its tables need inside its prepass (0.38 s of the STARK base's first PROVE SPLIT on the G1 traced run, with the card idle), and each device NTT twiddle size and staging pair at its first use, between two kernels. These entry points let a caller build the same state beforehand, off the prove: - `Domain::from_options`: `Domain::new` reads only the options from its AIR, so the same domain can be built before any AIR exists; `new` delegates. - `prover::warm_domain_and_twiddles`: fills the process-wide cache entry that `domain_and_twiddles` looks up, through the same function, so a warmed prove reads exactly what it would have built. - `gpu_lde::prewarm_device`, over `device::prewarm_twiddles` and `device::prewarm_staging_pairs`: the backend, every twiddle size up to a bound, and a number of staging pairs, built by the functions that build them on demand. Nothing calls them yet; no value a prove commits or reads changes.
…VM_OOD_COLUMNS_ON_CALLER) A round-3 OOD table is one or two rows high, and `Table::columns` transposes it with a rayon parallel iterator. The per-table drivers of `multi_prove` are plain OS threads, so each call is injected into the global pool and the driver waits for a worker, with its table's next device work unsubmitted. When the pool is busy with other work, that microsecond transpose waits behind it: on the STARK base's G1 traced run the host-only OOD absorb summed 5.16 s over split 9's tables and 2.73 s over split 11's, while the level-0 lead-in was verifying base epochs in the base's tail, against about 0.02 s in a quiet split. Three sites read these columns per table: the round-3 absorb and the two device DEEP dispatches (and the host DEEP loop). `LAMBDA_VM_OOD_COLUMNS_ON_CALLER=1` builds them with `Table::columns_serial`, the same per-column read on the calling thread. Unset or `0` keeps the pool (today's); anything else panics. The values and their order are those of `Table::columns`, so the transcript and the proof are the same bytes. Tests (`tests::host_schedule_tests`): the serial transpose equals the parallel one; the switch's parsing; and a whole multi-table proof, grinding off, is byte-identical with the switch on and off, with a counter showing the named arm ran. Reversing the serial column order fails the byte test. The file also tests the warm-up entry points of the previous commit: an options-built domain equals the AIR-built one, and a warmed cache entry is the one a lookup hits.
…AD_AHEAD) Before the card has anything to do, the serial head commits DECODE's precomputed columns on the host (RPX over a 2^22-row LDE, about a second on the block ELF, before epoch 0 executes), epoch 0's builder then initialises the device, and the first prove's prepass builds every domain and twiddle set. On the G1 traced run that is 2.44 s from run start to the first PROVE SPLIT, plus 0.38 s of prepass, all with the card idle. `LAMBDA_VM_BASE_HEAD_AHEAD=1` starts two helpers beside the producer: one initialises the device, commits DECODE there and then builds the device's per-size twiddles and four staging pairs; the other builds the host domains and twiddles of every trace size up to the epoch's. The epochs wait for the commitment where they first use it (each epoch's preparation, or its prove), and the base's observer hears it from the helper, before any epoch is prepared. Unset or `0` keeps the serial head (today's); anything else panics. Each helper step prints a `BASE HEAD:` stamp under `LAMBDA_VM_BASE_SPLIT`. The device commitment is `decode::compute_precomputed_commitment_device_or_host`: DECODE's five precomputed columns, row-major, through `lfm::commit::commit_group_device_or_host_with`, whose device root its device tests pin to the host's; `multi_prove` also rebuilds DECODE's precomputed tree at the first epoch and refuses the proof if its root differs. The host arm is the same interpolate, coset-evaluate and commit as the serial head's. Tests: the device-or-host root equals the host root in both leaf layouts, for a 13- and a 4,097-instruction program (the latter crosses the device's commit floor on a device build) and from an ELF; building the group column-major fails both. The head run ahead proves what the serial head proves (the shared comparison, now `assert_same_proved`, also used by the prep-ahead test), and the observer hears the serial head's commitment exactly once. The switch's parsing.
…TREE_TAIL_THREADS) The level-0 lead-in builds wrap prologues in the base's tail: each verifies its base epoch, fanning out over every thread of the global rayon pool while the base still proves its last epochs. The base's per-table drivers are plain OS threads, so each parallel iterator they start is injected into that same pool and waits for a worker. On the STARK base's G1 traced run the prologue spans hold 4.83 s of the base's 10.88 s of card idle, and the base's host-only OOD absorb, one such injection, summed 5.16 s over split 9's tables against about 0.02 s in a quiet split. `LFM_TREE_TAIL_THREADS=<n>` runs each prologue inside a rayon pool of n threads of its own, so it fans out over those and never queues ahead of the base's work in the global pool. The helper takes the lead-in's context before entering the pool, so no pool thread waits on the base. Unset, empty or `0` keeps the global pool (today's); anything but a count panics. The prologues compute the same programs and arenas either way; the tree's program ids are the check on a box run. Tests: the switch's parsing, and work run in a pool of its own fans out over that pool's threads and returns what it computed.
`LAMBDA_VM_GPU_DEVICE_ONLY_THRESHOLD` decides which tables keep a host copy of their LDEs, and those copies are downloaded inside each table's commit: on the STARK base's G1 traced run, 39.74 GB of retained-LDE downloads (8.08 s of driver time), about 80 % of it KECCAK_RND and ECDAS tables that sit below the 2^19-row default only by row count. A run that moves the threshold has to say so in its log, so the resolved envelope is printed once, where it is read: `[gpu] device-only envelope: LDE >= <rows> rows (<source>)`. No behaviour changes.
Two checks for a device build, so a host fallback cannot pass as the device: - the head helper's device stamp says whether the device took the warm-up (`BASE HEAD: device warm-up` or `... declined`), since a decline is silent by design (the prove then builds the same state on demand); - `the_head_decode_root_is_committed_on_the_device` (cuda): the 4,097- instruction DECODE map, 8,192 rows at blowup 2 and so above the device's commit floor, moves the device group counter and still gives the host root. The counter is process-wide, so the test is run on its own.
The ZfFormat::DEFAULT doc quoted a block timing for whir_grind=query from an earlier measurement. A measured number in a comment goes stale, so the doc now says what the lever does and why it loses no proven bits. The same reasoning in whir_chain's module header, GrindBits::query_only and WhirGrind gave the out-of-domain grind's reason as "it follows the out-of-domain point". That is half of it: the grind sits right before the batching challenge and does guard it. Dropping it costs nothing because the batching challenge has far more bits than the target without any grind. Comments only.
the_inner_node_verifies_two_leaf_nodes needs FAN_IN^2 = 4 epochs and took FIXTURE_EPOCH_LOG2 - 1. Since 8f9aef1 moved the fixture to the 48-cycle continuation-fixture at FIXTURE_EPOCH_LOG2 = 5, that is 16-cycle epochs and three of them, so this box-tier test has stopped at its own epoch-count assert in setup, before proving anything, under either wrap attestation. A pre-existing fixture drift, found by the gates at e413989. Two below the shared constant is safe by an assertion the suite already holds: the_fixture_guest_commits_in_an_intermediate_epoch keeps the guest above one shared epoch and within two, so 8-cycle epochs give at least five (six today), where 16-cycle ones give three or four.
…le v2) Job 202's NICE arms lost 9.65 s at STARK level 0 while their mechanism moved the right way (level-0 held time -5.4 s). The cause was the one host-phase pool shared by the six level-0 workers: every whole host phase was an injected job there, and a rayon thread blocked in a join runs injected jobs before it returns to its own (rayon-core 1.13 wait_until_cold). A thread waiting inside one wrap's reconstruct ran other wraps' whole phases nested on its stack, so the reconstructs that started first finished last (2-3 s became up to 22 s) and the card sat with no holder for 17 s of the level. host_phase now installs into a pool of the CALLING thread's own, built at its first host phase and dropped with the thread. A worker runs one phase at a time, so its pool only ever holds that phase's jobs; a caller that is already a rayon worker runs the phase inline. Each pool prints one line naming its caller and how many of its threads took the nice value, counted after every thread has started. Knob unset: inline, as before. Tests: two callers' phases never share a thread (with the shared-pool design as the control, which puts both on the same threads); a caller reuses its own pool; a rayon worker runs a phase inline; unset runs on the caller; a new pool reports every thread.
A table outside the device-only envelope downloads its whole main and aux LDE inside its commit, on the driver that would otherwise submit its next table. At the 2^19 default the STARK block's base downloaded 39.74 GB that way (8.08 s of driver time), most of it KECCAK_RND and ECDAS tables of 1,480 and 521 columns whose device paths already run: they sat outside the envelope by row count alone. At 2^16, with the barycentric floor (trace >= 2^14 rows) below it, the base downloads 7.27 GB and the block proves 1.40 s faster (FAST job 206, ds850-861, two arms each: wall -1.40 s, base -1.10 s; program ids unchanged, the root verified). `LAMBDA_VM_GPU_DEVICE_ONLY_THRESHOLD=524288` restores the 2^19 envelope. The default is process-wide, so the WHIR pipeline's recursion proofs (FRI STARKs through the same gate) take it too; that pipeline was not measured.
The head helpers (DECODE committed on the device, the first prove's domain, twiddle and staging state built beside epoch 0) become the default: `LAMBDA_VM_BASE_HEAD_AHEAD` unset, empty or `1` runs the head ahead, and `0` is the named opt-out that runs the serial head as before. FAST job 206 (ds850-861, two arms each against the serial head): the head 2.4 -> 1.0 s, the first prepass 0.37 -> 0.00 s, the base -2.00 s, the block -1.45 s; program ids unchanged, the root verified. The switch's test now pins the new reading (ahead unless exactly `0`); the equivalence test still proves both heads by parameter.
… by default The level-0 lead-in's prologues put parallel work on whatever pool they run in while the base still proves its last epochs, and the base's per-table drivers, plain OS threads, queue their own parallel iterators behind it in the global pool: the host-only OOD absorb summed 5.2-5.3 s over split 9's tables in all three G1 arms. In a pool of 16 threads of their own (FAST job 206, ds850-861, two arms each against the global pool) the base's worst split absorb fell to 0.04 s, the base by 2.35 s and the block by 2.20 s, level 0 +0.05 s; program ids unchanged. The pool now belongs to each `LeadIn`, sized by its caller: the STARK tree defaults to `STARK_TAIL_THREADS` (16); the WHIR tree keeps the global pool, its tail not having been measured in one. `LFM_TREE_TAIL_THREADS` overrides either (`0` is the global pool, the named opt-out; a count, a pool of that size), and the line naming the choice is printed where the lead-in starts. The rationale no longer says the prologues' verify fans out over the pool; what is established is that isolating them removed the stalls. Tests: the switch over each pipeline's default; a lead-in with a pool of 3 builds its prologues on that pool's threads and one without a pool does not (routing the prologue around the pool fails it).
…ries the_inner_node_verifies_two_leaf_nodes checked the composition property, that a node's published schema does not change with its level, as inner_layout.total() == leaf_layouts[0].total(). A node publishes its last child's output halves (emit_node_publishes), so that compare also required the FIRST leaf's carried output to be as long as the LAST leaf's: a fact about which epoch committed, not about the level. It held while neither leaf's last epoch committed. At the 8-cycle epochs the gate now runs at, the fixture commits in epoch 1, the first leaf's last, so that leaf publishes 146 words against the inner node's 144, under either STARK wrap attestation. Compare the words the two proofs published instead, the inner node against the last leaf, whose output halves it carries. The compare of two layouts built from the same count could not fail; the published counts can.
LFM_TREE_TAIL_THREADS now also takes `per-helper:<n>`: each lead-in helper builds its prologues in a rayon pool of n threads of its own. The default does not change (one shared pool of 16 for the STARK tree, the global pool for the WHIR tree), and `0` is still the global pool. Why: each helper installs a whole prologue, with joins inside, into its pool. A pool thread waiting in a join also takes injected jobs, so in a pool two helpers share it can run the other helper's whole prologue nested on its stack and finish its own only after that one; F-PREP measured that inversion on level 0. With one pool per helper, each pool has one caller and the inversion cannot happen. A per-helper pool of no threads is refused (rayon would read 0 as every core). A test checks that two helpers building at once run on two pools, each prologue on one pool of the per-helper width. Routing every helper to the first pool fails it.
Conflict in prover/src/zf_format.rs resolved by keeping both defaults: one_row=auto (the STARK pipeline's measured setting) and whir_grind=query (P2-W, which changes no STARK proof). The default banner is now cap=auto whir_cap=auto fri=dp one_row=auto whir_folds=first6 whir_stack=27 whir_grind=query.
…elper by default The STARK tree's default for LFM_TREE_TAIL_THREADS is now one pool of 8 threads per helper, the same 16 threads the shared pool had. LFM_TREE_TAIL_THREADS=16 is the named way back to one shared pool, and 0 is still the global pool. The WHIR tree keeps the global pool. A pool per helper has one caller, so a pool thread waiting inside one prologue can no longer run the other helper's whole prologue nested and finish its own late. FAST job 2115 (ds884-887, S P P S, two arms each) measured the per-helper pools against the shared pool: - wall +0.30 s, base -0.20 s, level 0 +0.70 s; - all inside the 0.8 s noise, and program ids unchanged. The only prologue it slows is the lead-in's last, which runs alone: 5.4-5.5 s on 8 threads against 3.9-4.1 s on the shared 16.
…by default STARK block ABBA: -4.15 s (EFFECTIVE). LAMBDA_VM_STARK_WRAP_FOLD=1 restores the in-guest fold and today's program ids byte for byte. Also repairs the box-tier inner-node test (a quarter-epoch fixture tree; the compare against the leaf it carries), whose failure was pre-existing at 8934b59.
…ad-in Head ahead (LAMBDA_VM_BASE_HEAD_AHEAD), the device-only envelope from LDE 2^16 (LAMBDA_VM_GPU_DEVICE_ONLY_THRESHOLD) and the tree's lead-in prologues in their own pools, 8 threads per helper (LFM_TREE_TAIL_THREADS). STARK block ABBA: -5.05 s (confirmed); per-helper pools NO EFFECT against one shared pool.
Both RPX implementations (prover lfm::rpo::Rpo256::mds, used by the block path's RpxStarkHash, and crypto::hash::rpx::mds) built each output lane with a core::array::from_fn closure. The closure's generic from_fn wrapper is placed in a codegen unit of rustc's choosing and is inlined into mds only when that unit happens to be mds's own. When it is not, every lane is an out-of-line call that recomputes (j - i) mod 12 with a 64-bit multiply per term: about a fifth more instructions per permutation. That is the two-speed host verify on the STARK tree. A Linux x86-64 cross build of the prover test crate at 8934b59, ef6d4be and 3fd644e shows mds inlined (2499 B, no calls) only at ef6d4be, the one FAST build, and twelve closure calls at the other two, the SLOW builds. Crypto's mds makes the twelve calls in its current partitioning too. Loops over a precomputed circulant compile the same way in every build. The permutation's values are unchanged: the RPO and RPX known-answer vectors, the two-implementation agreement test and a new test against the circulant definition all pass.
LAMBDA_VM_GAP_PREP_NICE now defaults to 10: the tree driver's reconstruct, emit and harvest and every LFM prove's execute and fill run on a pool of the calling thread's own, its threads at nice 10, so the proof holding the card keeps the CPU. LAMBDA_VM_GAP_PREP_NICE=0 is the opt-out and restores the schedule before the knob; 1..=19 picks another value. Measured at bbdac70 on FAST (F-PREP job 224, ABBA, one binary): the STARK tree took 1.55 s less (A 73.8 / 73.7 s, B 72.4 / 72.0 s). Level-0 card holds shrank 4.15 s; the card waits 2.49 s longer for the next wrap's host work. No wrap's reconstruct slowed past the stall guard (3.27 s against 6.26 s). Proofs are unchanged: a host phase moves where execute and fill run, not what they write (trace_identity_tests::execute_and_fill_on_the_host_phase_pool_are_byte_identical). The lead-in's prologues are untouched: they run in F-SIDLE's per-helper pools and never go through host_phase.
…GPU_COMPILED_CONSTRAINTS) The STARK quotient's constraint_composition_kernel is an interpreter. For every node of every row it: - loads the node; - decodes its operands; - reads them from, and writes the result to, a per-thread slot file in global memory. The slot file caps its grid at 65,536 threads, which is 0.30 waves on the 5090. G5's ncu of a base launch reads 77 % of L2, 30 % issue and 26 % occupancy: memory- and latency-bound on the interpreter's own traffic. The kernel takes 1.53 s of the STARK base's wall and 1.23 s of the recursion's (G1). This adds a straight-line twin of that kernel for every production program of at most ~4 k nodes: - the VM tables, the per-epoch local-to-global table and the LFM chips; - 36 kernels in all; - KECCAK_RND, ECDAS, ECSM and KECCAK stay interpreted. Each node becomes the call eval_program_row makes for its op, operand kinds and result class, over local variables, with no slot file and a grid that fills the card. The transition sum adds the roots in root order, and the boundary tail is the interpreter's own. Goldilocks values on the device are non-canonical u64s, so this exact mirroring is what makes H bit-identical. stark::constraint_ir::codegen emits the kernels, keyed by a structural hash of the lowered program. The prover crate's tests::compiled_constraints generates crypto/math-cuda/kernels/constraint_compiled.cu and its key table, and fails when the committed source is stale. math-cuda loads the module on first use. A program with no kernel, or a module that cannot load, runs the interpreter. LAMBDA_VM_GPU_COMPILED_CONSTRAINTS=1 turns it on. Unset or 0 keeps the interpreter, the default. Anything else stops the run. Evidence: - The host build of both CUDA sources (crypto/math-cuda/kernels/tools/ccomp_host_check.cpp) agrees limb for limb on all 36 programs, over random full-range inputs. - Three deliberate generator faults each fail every one of the 72 runs. - The freshness test fails on a one-character edit of the generated source. - Device tests compare the two kernels on the card and prove a program both ways, byte for byte.
…aces, with a mutation control `the_compiled_kernels_prove_the_same_bytes` proved add.elf three times through `prove_with_options`, and its control (two interpreter proofs) failed on the card: two builds of one program's traces differ. The LT, BRANCH, MUL, DVRM, EQ and BYTEWISE builders deduplicate through a std HashMap and lay rows out in its iteration order. On the laptop, two builds of add.elf's traces give different LT rows and different proof bytes. The test now builds the traces once and proves copies of them (`FixedTraces`, `prove_with_options`' prove step on a clone), at grinding 0: - the control: two interpreter proofs, byte-equal; - the compiled proof: byte-equal to them, verified, with compiled compositions counted; - the mutation control: BITWISE's kernel swapped for a mutant whose last root is off by one. The proof must change and fail to verify, or the prover must refuse it, and the mutant must have run. `one_set_of_traces_proves_the_same_bytes` runs the control alone on the path the build proves on (ignored: two proves at blowup 4). Supporting changes: - `Traces` derives Clone in this crate's tests only. - codegen: `mutant_composition_kernel` (the kernel plus one wrong statement) and `mutant_kernel_name`. - The generated source carries BITWISE's mutant last. It is not in the key table, so no program runs it. - gpu_interp: `substitute_compiled_kernel` and a call counter, compiled only under stark's `test-utils`, launch a named kernel in another's place. The 36 production kernels and the key table are unchanged.
… streaming verifier derivation) into logup/s2 No conflict. Under the default pair policy every program id this head pins is unchanged (the_compact_program_form_keeps_a_small_trees_ids passes on the merge), as are the 44 production programs of the pair golden, the compiled kernels and the stark goldens.
…rates The calibration trimmed the pool and read the card's free memory while frees were still queued on the base's streams: a queued free holds its block until its stream reaches it, and the trim cannot hand back what the pool has not been given. At the 1x level-0 arming that left 11.6-12.8 GiB used after the trim, of which 0.33 GiB was live; draining first leaves 1.3-1.5 GiB (FAST 479a/b readout runs), so the budget is the configured cap instead of 15-16 GiB. The drain runs only when the gate calibrates (nothing admitted, no claim in force), between sibling levels.
… into sched/l2p at step 3 (73e3d64) The landing train for lever 2: the shared VRAM gate (carried residents behind per-prove claims, the pool releasing only while armed, the arming drained first) on #1013's current head. No conflicts; #1013's changes do not touch the gate, the admission or the device-set model the estimates come from.
pin_shared_vram_gate is process-wide; under cargo test's threads a pin could hand a concurrent resident_carry_tests prove the shared gate and move its account readings. One test-only lock orders them (nextest's process-per-test runs were never exposed).
The block tree's sibling proofs now share the card by bytes unless LAMBDA_VM_SHARED_VRAM_GATE=0 restores the exclusive card permit. The gate still acts only while a caller arms it for concurrent proofs (device_permit::arm with more than one worker), so the base and every single-prove path are unchanged, and proof bytes do not move: the gate changes admission order only (one top id across every A/B run, FAST 477-479). FAST 479 P8 against the exclusive card: 1x recursion -0.59 s (t -17.2), whole -0.51 (t -6.9), base -0.02, VRAM <= 26.1 GiB. The permit tests that check the exclusive card pin the gate off; without the pin the falsifier test now fails, which is the flip taking effect.
A block plan absorbs the ELF's digest into every leaf, and the fixture ELF's bytes depend on the clang that assembled it: the laptop's (digest 3b39e219…) and FAST's (f5e120d0…) differ, so a single pin set holds on one machine only (FAST 670). The two pinned-id tests now look up their ids by the plan's ELF digest, with the laptop's pins (recorded at 541f4bd) and FAST's (recorded at fe1fb16 in FAST 672, where the compact form's ids matched them on the device and the host). An ELF with no pins refuses by name before any program is derived; pinned_tree_ids_refuse_an_elf_they_do_not_know checks that.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…14fea1 crypto/stark/src/narrow.rs and crypto/stark/src/spill.rs are #1013's files, byte-identical to mem/stark-f1 @ e5214fea1 (i-mem3's F1: the page-backed Bytes, the streaming Digester, the fused writer, over the S1/S2 store). #1013 is their one source of truth: #1014 never edits them, re-imports any change from there, and they merge as one when both PRs reach main. Wiring only: the two modules in lib.rs (dead code allowed there, since some parts serve #1013's prover alone), math's page-bytes feature, and libc unconditional as the store needs it (it was behind disk-spill). #1013's prover glue (TraceTable::spill_main) and its spill tests are not imported; #1014's adapter and tests follow.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…ault off) LAMBDA_VM_BLOCK_SPILL=always (BlockOptions::spill, BlockSpillPolicy) opens #1013's spill store for the prove. Phase A hands each committed group's packed tables to it as the group is installed, past the first two groups (phase B reads them before a read-back could land) and never past the writers' queue: the committer is the store's one producer and checks the room first, so it never waits on the writers. Phase B reads the spilled tables back in group order, two groups' bytes ahead of their uploads, restores each group and the next one before their uploads, and lets a finished group's packed columns go. While the spill is on, the packers put the packed bytes in pages of their own (the card's pack and b2's finish pack), which the writers take whole. The adapter between NarrowColumns and the store's NarrowMain moves the parts and never copies the bytes (a test checks the address). A read that fails or does not match its digest refuses the prove with MlError::SpillFailed, naming the table. BLOCK SPILL reports the store's counters and the read-back. Off by default: with the spill off nothing changes on any path (the heap, the same allocations, no store). Tests: a spilled block proves the same proof bytes as a held one under the deterministic grind, each slot read back once; a byte flipped in a spilled table on disk is refused.
At the median the setup line's "unnamed" read 6.31 GiB in every run and arm of BIG 110: deterministic, so sized by the program or its pages, not the run. The line now reads the heap before VmAirs::new as well and names what the AIRs and the AIR/trace pairing add. Measurement only (memlog).
…unted apart in BLOCK MEM
The two spill measurement rounds (BIG 115, 116) read the writer's time by
step; this keeps that split for whoever measures the spill next.
- SpillStats splits the writers' seconds into digest, copy (the aligned
copy under O_DIRECT) and pwrite; the rest is the free and the locks.
The BLOCK SPILL line prints them.
- SpilledMain::is_resident. The BLOCK MEM traces/setup lines count
spilled traces apart, and their bytes in the heap only while a slot
holds them; they were counted at 8 bytes a cell ("traces 228.87" at the
median with the spill on).
Test: a store's step seconds add up to no more than its writer seconds,
the copy appears exactly under O_DIRECT, and written slots leave memory.
With LAMBDA_VM_BLOCK_SPILL unset the block now runs `auto`: a committed packed trace is spilled only once the host's bytes (the larger of VmHWM and the cgroup's memory.current), the reserve for what the block still needs and the trace would pass the target (memory.max, or MemTotal, less 10 GiB). A block that fits spills nothing; a block above the target spills what would not fit instead of running out of memory. `off` keeps every trace, as the default did before. No proof byte depends on the policy: the words read back are digest-checked. Tests: unset reads as auto; the stream under auto precommits the same instances and builds the resident traces.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
… reserve LAMBDA_VM_BLOCK_SPILL now reads as #1013's (prover/src/block.rs SpillPolicy … spill_decision @ edddc68): auto (and unset) | off | always | <GiB> resident budget. auto spills a committed table once the host's bytes (the larger of VmHWM and the cgroup's memory.current), the reserve and the table would pass the target (LAMBDA_VM_BLOCK_SPILL_TARGET_GIB, else the cgroup's memory.max less 10 GiB, else MemTotal less 10 GiB). Phase A decides per table and never past the writers' queue; groups 0 and 1 count but never spill. The reserve is #1014's own, inferred from the median block (BIG 565): the finish's p5 transient (25.5 GiB) and the tree's leaf programs alive through phase B (16.3 GiB), about 1.25 GiB per total G cells or 1.65 per G cells committed so far, plus 6 GiB for phase B's bump and the read-back window. BlockSpill carries the policy's choice (SpillWanted) and the kept bytes and committed cells it reads. A block that fits spills nothing; the store opens for every policy but off. Tests: the knob and the decision (unit), a budget the block fits in spills nothing and one of zero spills (block).
…Total i-m4b found the gap: the target read only cgroup v2's memory.max. On a v1 host with a limit below MemTotal (FAST: 57.53 GiB against 59.93) it fell through to MemTotal less 10 GiB and aimed 2.4 GiB too high there. - cgroup_memory reads v2's file under the unified hierarchy, else v1's under the memory controller (memory.limit_in_bytes for the limit, memory.usage_in_bytes for auto's host charge), each at the process's cgroup path and then at the hierarchy's root, which is what a container without a cgroup namespace sees (FAST: 12:memory:/docker/<id>, the limit at /sys/fs/cgroup/memory/memory.limit_in_bytes). - The target is min(cgroup limit, MemTotal) less 10 GiB (spill_target_from): v1's unlimited sentinel gives MemTotal's, and a v2 limit above MemTotal no longer wins. Tests: fake cgroup trees for v2 at its path, v2 `max` falling to v1, the FAST layout, v1 at its path, a shared v1 hierarchy, and none; the target rule with v1's sentinel, a cap at MemTotal, and neither. Dropping the v1 read or the minimum fails them.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…@ 278e6a8 The rule #1014 copies from #1013 read only cgroup v2's memory.max and memory.current. On a v1 host (FAST) the target fell through to MemTotal less 10 GiB, and the host's charge was not read at all. #1013 fixed both in 278e6a8; this re-copies spill_target_bytes, spill_target_from, cgroup_memory and host_bytes_now from that sha, with its two tests (the cgroup files from fake trees, and the target rule), cited as before. The only changes from #1013's text are the citations and the tests' temporary directory, which gets a name of its own so the two copies of the test never share one if they ever run in one process.
A carrying prove's claim held its tables' resident bounds plus the largest table's whole fused set, whose own main set was already among the residents. At the median a leaf claimed 9.4-10 GiB for 5.4-6 GiB of residents, and level 0 still waited on claims (Σ 6-9 s a run, BIG 473 A). LAMBDA_VM_SHARED_GATE_CLAIMS=tight (off: today's claims) claims the resident bounds R plus a headroom H = max_i max(fused set_i - resident_i, scratch_i): every carry is at most its bound, so in Round 1 a table asks at most its bound plus scratch beyond what the prove holds, and in its fused task at most its set less its carry beside the other tables' carries - both inside R + H whatever the commits leave resident. The claim book admits while Σ R + max H <= budget over the claims in force: in a state where every prove waits, the gate holds at most Σ R and every request is at most max H, so every waiting prove fits, and a finished table, a settle or a departure only lowers the left side. Keeping the smallest H instead is not safe; a test shows the wedge. A leaf's tight claim is ~7.4 GiB and three leaves share one headroom (18.7 GiB of the 23.44 budget). Tests: two tight claims whose top-ups would wedge without the headroom; the largest-headroom book against the smallest (X, Y, Z); a randomized run of 4 threads x 6 proves x 4 rounds with host fallbacks (never wedged, Σ R + max H <= B with two claims in force); the wiring test (proof bytes equal, headroom covers every top-up, settled = carried + headroom, below today's claim). Admission without the headroom, or with the smallest, fails them.
auto took the host's bytes as max(VmHWM, the cgroup's charge), and the charge counts page cache: v2 memory.current and v1 usage_in_bytes alike. On a cache-heavy box it spilled blocks that fit. FAST at 17:28Z charged 40.58 GB with 15.18 GB of it inactive file pages. At BIG 118's p75 the spill-off control fit at 116.1 GiB by the kernel giving back 8 GiB of cache, while auto had spilled 29 GiB. - HostReading: VmHWM, the charge, and memory.stat's inactive file pages (v2 inactive_file, v1 total_inactive_file, read by their exact key). The host's bytes are max(VmHWM, charge - inactive_file): the cgroup's working set, the kubelet's measure. Inactive file pages are what the kernel reclaims first; active file pages stay counted, and the target's 10 GiB margin covers them. - cgroup_memory reads a whole-file number or one memory.stat key. - The BLOCK SPILL line prints the largest reading auto decided on (VmHWM, charge, inactive file). Tests: fake memory.stat trees for v2 and v1 (the hierarchical key, an exact match past a decoy), and the working-set arithmetic. Not subtracting the inactive pages, or matching the key by suffix, fails them.
The stream's byte budgets (ops 4 GiB, generated 2 GiB) bound the work piled up behind a slow spill writer, but they applied whenever a store was open. Under the default auto that is every block: at the p75 they bound with or without spilling (ops waited 10-12 s, generated 94-121 generator-s, against no-spill queues of 4.78 / 3.36 GiB that never waited) and cost phase A about 3.6 s. The budgets now arm (QueueRoom's Arming, from Spill) only where a spill is plausible: always under `always` or a budget, never with the spill off, and under auto once the host's working set plus the reserve reaches 85 % of the target (SPILL_ARM_SHARE). Once armed they stay armed. Unarmed, the queues run as with the spill off, so blocks far below the target, the 1x block and the median among them, pay nothing. The BLOCK QUEUE line says when the budgets armed, or that they never did. Tests: an unarmed budget admits beside a full queue and binds once armed; the arming rule per policy and at 85 %. Arming that is ignored, or a zero share, fails them.
A carrying prove now claims its tables' resident bounds plus the shared headroom by default; LAMBDA_VM_SHARED_GATE_CLAIMS=whole restores the first form (the whole bound held, no headroom). BIG 474 at the median, 5 + 5 on one binary: recursion -1.25 s (t -5.4), level-0 claim waits from 8-10 s a run to 0, VRAM <= 28.7 GiB, host <= 85 GiB, one top id; FAST 479-p12 at 1x: recursion +0.00, no regression. Scheduling only: claims order admission, so proof bytes do not move. A test that forces the carry picks the claim's form itself, whatever the knob says.
… as a knob (default off) Under the posture's never-purge jemalloc a phase rarely reuses the pages the phase before it freed (they sit in other threads' arenas or in other size classes), so a block's RSS ratchets phase by phase. At the spill-on median (BIG 584) the live peaks are flat at ≈ 36 GiB across the base, the tree and the verifier, while VmRSS climbs 52.75 → 56.56 → 60.28 GiB. Decay 0 returns pages continuously, at +54 s of base. alloc_purge::purge_point(point) runs one arena.<all>.purge (jemalloc's MALLCTL_ARENAS_ALL) when LAMBDA_VM_ALLOC_PURGE names the point (all, or a list split on commas or dots). It prints an ALLOC PURGE line with its wall time and jemalloc's resident bytes before and after. It purges only in the lib's test builds, whose global allocator is jemalloc; elsewhere it does nothing, so pipeline code can call it unconditionally. The block calls it at three points: phase-a (block.rs, phase A done), base (the whole-block harness, the base proved) and tree (the top proved, before the block verifier). tikv-jemalloc-sys joins the dev-dependencies, without default features, for the no-value control. Unit tests: the knob's names; a purge returns 512 freed 1 MiB buffers (laptop, posture: resident 525 → 4 MiB in 4 ms; without the posture the test refuses and says why).
i-tree's knob purged only where LAMBDA_VM_ALLOC_PURGE named a point. At the p90 block 25481021 the two purges are what make it fit: without them level 0 was OOM-killed at 119.79 GiB (BIG 120); with phase-a and base purged, level 0 started at 55.6 GiB and the whole block verified at 116.6 GiB (BIG 1215). The full-gas block went 120.22 -> 111.3 GiB and 25 s faster. - `auto`, the new default (unset, empty or `auto`): purge at phase-a and base once the block's memory is short. The spill's queue budgets arming (host + reserve >= 85 % of the target) calls note_memory_pressure(); each block clears it at its start. A block that fits never arms, so it pays nothing, and the line says "skipped (auto, no memory pressure)". - `tree` is not an auto point: it trims the block verifier's start, not the proof's peak. `off` purges nowhere; `all` and point lists purge as before, regardless of pressure. - The purge still runs only in the lib's test builds (jemalloc as the global allocator); the CLI's own allocator has no purge yet, a landing item beside the production whole-block driver. Test: auto purges at phase-a and base only under pressure and says it skipped otherwise, never at tree; off and all as named. Ignoring the pressure fails it.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…s it from plain threads BIG 569's arm N parked in the W3 tree: its leaves' artifacts were built as rayon jobs that each held the card around a build that uses rayon, and a holder that waits inside rayon runs its pool's queued jobs on its own thread, so a sibling build there was a second hold on one thread (i-m4b's diagnosis and repro). Which run trips it is the scheduler's choice; the node pipeline's pool tasks, which built their node's artifacts too, had the same shape. - The armed permit refuses a hold on a rayon worker, before it takes the card (the repro, fe3dd12, now asserts that refusal; a plain thread holding the card while rayon builds under it, and a second plain holder waiting, are tested too). - The W3 tree builds its leaves' artifacts one after another on their own thread (each build held the card throughout, so the rayon jobs bought no overlap). - The node pipeline emits a level's programs on its pool and builds each one's artifacts on the builder's thread as it arrives (#1013's split): build_levels takes an emit (pool) and a finish (this thread); the out-of-order test checks every finish ran off the rayon workers (finishing on a worker, or publishing in arrival order, fails it). - The harness disarms the permit before its off-the-clock verify, whose WhirBlockPlan::programs builds artifacts as rayon jobs.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…s it from plain threads BIG 569's arm N parked in the W3 tree: its leaves' artifacts were built as rayon jobs that each held the card around a build that uses rayon, and a holder that waits inside rayon runs its pool's queued jobs on its own thread, so a sibling build there was a second hold on one thread (i-m4b's diagnosis and repro). Which run trips it is the scheduler's choice; the node pipeline's pool tasks, which built their node's artifacts too, had the same shape. - The armed permit refuses a hold on a rayon worker, before it takes the card (the repro, fe3dd12, now asserts that refusal; a plain thread holding the card while rayon builds under it, and a second plain holder waiting, are tested too). - The W3 tree builds its leaves' artifacts one after another on their own thread (each build held the card throughout, so the rayon jobs bought no overlap). - The node pipeline emits a level's programs on its pool and builds each one's artifacts on the builder's thread as it arrives (#1013's split): build_levels takes an emit (pool) and a finish (this thread); the out-of-order test checks every finish ran off the rayon workers (finishing on a worker, or publishing in arrival order, fails it). - The harness disarms the permit before its off-the-clock verify, whose WhirBlockPlan::programs builds artifacts as rayon jobs.
arm_running_total(margin) puts the pool in the armed gate's release posture and returns the budget arm_shared_vram_gate would calibrate (card free after a drain and a pool trim, less the margin, capped at the configured budget), without arming the shared gate; disarm_running_total puts the posture back unless the shared gate is armed by then. For the block tree derivation's own byte gate, which arms only while the shared gate is not.
BIG 123's p90 block (110 leaves) proved and verified, then the post-run block verifier ran the card out of memory: with the shared VRAM gate unarmed, each artifact commit was checked only alone, and a level's builds in parallel (each with up to eleven device sets in flight) asked for more than the card has. derive_top now arms a byte gate (lfm::derive_gate) while the shared gate is not armed: a running total of the commits' device sets, calibrated like the shared gate (card free after a drain and a trim, less 2 GiB, capped at the configured budget), with a set over the budget running alone. A commit holds its bytes around its device dispatch alone, which forks nothing, so no permit is held across a rayon wait; the build's fork (map_maybe_parallel) and the gate's admit both assert it. The level keeps its host parallelism. LAMBDA_VM_BLOCK_DERIVE_GATE: auto (default), off (no total, as before), or a budget in GiB. LAMBDA_VM_BLOCK_DERIVE_BUILDS=n now runs a level in windows of n instead of holding a count place across each build inside a rayon job, which a job stolen at the build's fork could ask for again on the same thread. Scheduling only: every artifact is a pure function of its program, so the derived top cannot change.
…n the card The whole-block harness reads the cold verifier's own device high-water from the pool (a drain, a trim and a restarted high-water first), not the card, and prints the derive gate's summary beside it. the_derive_gate_bounds_the_derivation_on_the_card (box tier, cuda): the six-leaf fixture plan derived open, then under half the open peak: the gated running total stays within its budget and waits, the top is the open one's, and a permit held across a build's forks is refused.
A thread that holds a derive permit and asks again now passes straight through with no bytes of its own instead of asserting: it cannot park the thread, at worst it over-admits, and the gate counts it as a re-entry. A build's fork (map_maybe_parallel) counts a permit held across it on that permit's gate instead of asserting. Both counts are in the gate's summary line and the tests read them, in debug and release alike; the prover's release path no longer has a panic for either state.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…cf253d2) The host measure #1014 copies from #1013 took max(VmHWM, the cgroup's charge), and the charge counts page cache. FAST charged 40.58 GB at 17:28Z with 15.18 GB of it inactive file pages, which would have made auto spill the last groups of a 1x block that fits. This re-copies #1013's measure from its landed head 035aef5 (commit cf253d2): HostReading reads VmHWM, the charge and memory.stat's inactive file pages (v2 inactive_file, v1 total_inactive_file, by their exact key), and the host's bytes are max(VmHWM, charge - inactive_file). CgroupValue lets cgroup_memory read a whole-file number or one memory.stat key. The BLOCK SPILL line ends with the largest reading auto decided on, as #1013's does. With #1013's two tests; the text differs only in citations and the test's temporary directory. The rest of #1013's train (lazy queue budgets, the allocator purge) is not copied: #1014 has no queue budgets.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One STARK proof per block, with no epochs (the prove-and-retire / VADCOP shape). This is a second prover next to #1009's epoch-based one. #1009 stays the reference and this branch does not touch it. The branch starts from #1009's head
cc411aa2c, so the diff against main includes #1009. Compare againstcc411aa2cto see only this work.Whole block: 31.78 s (FAST 456 at the head
08ecc4310, mean of 3 default arms), against the epoch tree's 60.15 s (#1009; FAST 389's same-binary reference, not re-run since). Base 21.28 s, recursion 10.25 s (level 0 5.18 · interior 4.88).Head
035aef5d6(10-03): the p90 and full-gas blocks prove end to end at the defaults on BIG (128 GiB host, 120.69 GiB cgroup), and the block verifier passes on both. p90 25481021 (52.0 M gas, 612.3 M cycles, 20.1× the bench block): whole 590.95 s at 115.79 GiB (BIG 591). Full gas 25431071 (60.0 M gas, 602.9 M cycles): whole 541.51 s at 110.62 GiB (BIG 125). Before this head,autoalone ran out of memory in level 0 on the p90 (BIG 120). Prover and harness only: digestac406fc8…is unchanged (FAST 515). At 1× nothing arms or purges (default against spill off, base −0.065 s, t −0.6, 8 + 8; FAST 515 + 516). Four changes:autocounts the cgroup's working set (the charge less its inactive file pages), not its page cache.autofell from +3.6 to +1.12 s (BIG 118 → BIG 122).auto, once the budgets arm). Blocks that fit pay nothing.LAMBDA_VM_ALLOC_PURGE=off|all|<points>. It acts in the lib's test builds (the harness's jemalloc); the production CLI has no purge yet.Head
00871b913(10-03): under the shared VRAM gate, a carrying proof now claims its resident bytes plus one headroom shared with the claims in force. These tight claims are on by default;LAMBDA_VM_SHARED_GATE_CLAIMS=wholerestores the first form. This is scheduling only: digestac406fc8…and top93aa7097…are unchanged. Median block (BIG 474, 5 + 5): recursion −1.25 s (t −5.4), with level-0 claim waits falling from 8–10 s a run to 0. Whole −0.05 s: the base moved +1.18 s, which is scatter. At 1× there is no regression (FAST landing gate: recursion −0.01 s). See Recursion: sibling proofs share the card.Head
edddc6873(10-03): the trace spill store is on by default asauto. A block that fits spills nothing. A block above the host's limit spills the committed traces that would not fit, instead of running out of memory.LAMBDA_VM_BLOCK_SPILL=offkeeps every trace. This is prover-only: digestac406fc8…is unchanged. BIG 117 at the median block: the default spilled 0 B, with phase-A end −0.97 s against off; forced to a 70 GiB target it spilled 30.1 GiB, cost +1.43 s and verified. The same head adds the spill writer's step timings and splits the AIRs out of BLOCK MEM's unnamed heap; both are diagnostics only. See Disk spill: auto by default.Head
a93018974(10-03): the block tree's sibling proofs now share the card through one VRAM gate, on by default (LAMBDA_VM_SHARED_VRAM_GATE=0restores the exclusive card permit). This is scheduling only: digestac406fc8…and top93aa7097…are unchanged. Against564dd02bb(FAST 871): 1× recursion −0.68 s, whole −0.48 s. Median recursion −4.72 s (BIG 471, measured before the arming drained the device); at this head −4.24 s (BIG 472). See Recursion: sibling proofs share the card.Head
564dd02bb(10-03): adds LogUp k4 as an opt-in (LAMBDA_VM_ZF_LOGUP=k4; defaultpair= the same bytes; under k4, median phase B −9.56 s and 1× whole −1.04 s).Head
8f2ce57d5(10-03): three block-tree changes, all with the same proof bytes and program ids (FAST 673 merge gate green: 1× top93aa7097…, pinned ids unchanged; median top6d85454d…, BIG 586).NOEPOCH_TREE_EMIT_LATE=auto): when leaf 0 puts the other leaves' programs at ≥ 8 GiB, they are emitted in phase B onto pages the retired traces freed; smaller trees emit at the shape as before. Median, spill off (BIG 585): base-phase VmRSS −10.37 GiB, whole-run peak 98.99 → 86.49 GiB, recursion −0.31 s, whole −2.15 s; inert with spill on.NOEPOCH_TREE_EMIT_LATE=offrestores the old behaviour, and A/Bs against it must set that.verify_block_tree): each program and its artifacts are dropped once its child is derived. Verifier alone in a fresh process at the median (BIG 587): peak 32.33 → 22.89 GiB, derive −1.48 s.LAMBDA_VM_BLOCK_DERIVE_HOLD=levelrestores the old hold.Note on drivers: the whole-block numbers here come from the test harness's tree driver (
block_tree_pipeline,cfg(test)). The library has the base prover (block::prove_block), the per-node program emitters (BlockTreePlan::leaf_program/node_program) and the production verifier (verify_block_tree), but no production driver that chains them into a whole-block proof yet; that is a landing item.Head
dacbd5b4e(10-03): KECCAK_RND, ECDAS and KECCAK compose on a bounded-slot interpreter by default (LAMBDA_VM_GPU_INTERP_SI=0restores the slot-file interpreter). Prover-only: digestac406fc8…unchanged, FAST 784's landing gate green (1× default − off −0.05 s).Constraint composition on a bounded-slot interpreter (prover-only; proof bytes unchanged). The three programs too large for the compiled kernels (KECCAK_RND, ECDAS, KECCAK) now compose on a bounded-slot GPU interpreter. The slot-file interpreter it replaces (from OpenVM v1) keeps every value of a row in a per-thread file in global memory: 4,350 words for KECCAK_RND, 2.1 GiB per composition, not counted by the VRAM gate. The new lowering evaluates each constraint's cone on demand in constraint-index order, reads trace cells as operands, and fits a row in 16–128 words of shared memory or a local array, recomputing past the budget; each program gets its measured best shape. On one binary, the three programs compose in 0.58–0.63× the time. At the median block (25475471, 3 + 3) the deciding figure is card work −0.43 s (t −1.9), with the slot-file scratch 80 GiB a run → 0 and the phase-B device peak −1.9 GiB (t −3.4); base −1.03 s is not significant (t −1.0); at 1× base −0.06 s. The digest is unchanged (
ac406fc8…);LAMBDA_VM_GPU_INTERP_SI=0restores the slot-file interpreter. Running all 39 programs on it is slower (1.25–2.1× the compiled kernels, +0.16 s at 1×), so the compiled kernels stay for the other 36.Head
d740eb5d5(10-03): the block tree emits each node's program as soon as its children land, instead of level by level (NOEPOCH_TREE_NODE_EMIT=levelrestores the old order). This is scheduling only: the same top program id in all 14 arms.NOEPOCH_TREE_EMIT_WINDOW; −12.9 GiB median base peak for +6.2 s recursion) and a cap on concurrent derivation builds (LAMBDA_VM_BLOCK_DERIVE_BUILDS).Head
86e71de77(10-03). Three landings sincef805464f6:aa726ce2b: block fan-in 4 (Mauro approved 10-02; see (f)).5e4176961: KECCAK and ECSM chunked, and tree partition rule v2, a format change. See the format section below.7ef261724+86e71de77: producer and memory work, all prover-only with identical proof bytes:LAMBDA_VM_BLOCK_GENERATORS=6restores six);LAMBDA_VM_BLOCK_SPILL=alwaysturns it on).The median block 25475471 on BIG (job 112, Zen 3 host):
ac406fc8…unchanged.Disk spill (BIG 111, base only):
The walker levers with eight generators end the median's phase A 3.45 s sooner (BIG 468) and cost +0.14 s on the 1× base (FAST 470).
Head
f805464f6: two more producer and memory changes, both prover-only with identical proof bytes:6cb6cd557, smaller per-step records: the CPU op shrinks from 128 to 96 B (arg2, the branch decision and the ECALL kind are derived) and MEMW_A ops become 48-byte aligned rows. FAST 612 at 1×: base −0.38 s, digest equal. BIG 424 at the median: phase-A end −6.0 s.f805464f6: LT ops kept as segments rather than concatenated, and the packed KECCAK_RND / LT finish builds back on by default. BIG 107 picked them over capping wide KECCAK_RND chunks: at the median they save about 14 GiB and end phase A 2.6 s sooner.BIG 108 at this head: 1× verified at 15.1–16.2 GiB with digest
ac406fc8…. The median block verifies at 81.8 / 82.4 GiB with base 212.4 / 216.0 s. Phase B scatters run to run by about 2.7 s (sd) at the median.Head
330057f2a: the packed KECCAK_RND / LT finish builds are now off by default (LAMBDA_VM_BLOCK_PACKED_BUILD=1turns them on), and the memory knobs that had no effect are removed. Two one-binary A/Bs at the median (BIG 104, 105): packed builds save 13–15 GiB of host peak but cost phase B ≈ 3–6 s. BIG 105 at this head: 1× verified at 18.32 GiB, digestac406fc8…; median block 96.2 GiB, base 218–224 s on BIG.Head
161b82415: six generator threads by default (LAMBDA_VM_BLOCK_GENERATORS=0keeps the committers generating). With the faster walk the three committers, which also generated every chunk, fell behind: at the median block 32.8 GiB of op lists waited and phase A ended 19 s after the finish. Generators build and pack each chunk; the committers only commit. KECCAK_RND and LT in the finish are built packed a block at a time (LAMBDA_VM_BLOCK_PACKED_BUILD=0builds them wide). BIG 103/104 on the median block 25475471: base 231.8 → 222–226 s, host peak 96.1 → 82.7–83.0 GiB; 1× base flat; digestac406fc8…unchanged. Within one binary, packed builds cost phase B ≈ 6 s at the median against wide builds and save 14.75 GiB; see the head above for the current default.Head
3be24abab: narrow storage on by default (merged froma2cd24207). Main-trace cells are stored at about 2 B each instead of 8 and widened on the card;LAMBDA_VM_BLOCK_NARROW=0keeps them wide. BIG 100 (A/B on one binary, the 128 GiB box): proof digests packed = wide =ac406fc8…; 1× host peak 31.2 → 17.7–19.0 GiB; 1× base −1.29 s on that box. A median mainnet block (25475471, 9.78× this one, 941 sub-proofs) proves and verifies within 128 GiB: host peak 106.87 GiB, base 259 s on BIG's Zen 3 host (BIG 100; that binary predates the E1+E4 producer changes below). BIG 102 re-gated the merged head: 1× verified, digestac406fc8…, peak 19.29 GiB.Head
2c440b8dc: base 19.61 s (FAST 602, mean of 2). It adds two producer changes to19fe5e9e5, whose memory changes are in Memory: a lean walk (the walker builds only what the tables need) and the executor's guest memory in 64 KiB pages instead of a hash map of words. A B B A on one binary: base 20.38 → 19.61 s (Δ −0.77, no overlap); proof bytes equal underLAMBDA_VM_FIXED_TRACE_HASH=1+LAMBDA_VM_DETERMINISTIC_GRIND(digestac406fc8…in both arms); host peak flat. At a median-size block (25475471, harness, windows only) the producer's windows phase falls from 42.4 to 25.2 s and the executor from 108 to 11 ns per cycle.LAMBDA_VM_WALK_LEAN=0andLAMBDA_VM_EXEC_MEMORY=wordsrestore the old paths. The whole tree has not been re-run at this head.The whole block, epoch tree vs no-epoch tree (FAST 389, one binary at
e4ff3f8fb, means of 2 arms)Block 25368371. Both trees run on one test binary, and its md5 is checked after every arm. The epoch arm is #1009's record harness, unchanged.
a0954fd1bthe lists are the plan's: D-NOEPOCH §12.2's rule over the closed-form costs, a pure function of the block's shape. On this block the rule gives other lists than §12.2's, which came from legmodel.py's costs, but the leaf tiers are the same. FAST 392 atcda620461(means of 2): level 0 6.21 s, interior 7.19 s, harvest 2.03 s, whole 40.38 s, which is 389's numbers. The verifier's own derivation of the top program takes 8.35 s; that is outside the whole, and the derived top equals the proved top. A tree over another partition, a skipped instance or a duplicated instance is refused at the final check (fixture gate).The plan's gate and the harvest lever (block 25368371, means of 2 arms, 389's posture)
cda620461(the plan)a684f4f3f(+ ELF constants beside the base)a1a3a2c22(+ F1 follow-ups)e7406d6b5(+ rebuilt windowed builder)6b86432a2(+ verifier derive)verify_block_tree, no proof read, outside the whole)e7406d6b5and FAST 397's6b86432a2are both on this branch;6b86432a2is the head.-x) ofnoepoch/windowed-builder6e99c74..1d24a2c, not a merge of that branch (90 WHIR commits, 6 conflicting files), plus ourparallelgating, which compiles the recursion ELFs (5fe8316, now byte-identical to the builder branch's bcccdef).verify_block_tree_with(elf, &ElfConstants, …)lets a consumer compute the per-ELF constants once (3.05 s cold), and each node level is now emitted and built in parallel (level 1: 1.31 s wall for 5.0 s of work). Split at 397: constants 3.05 · plan 0.16 · leaves 2.17 · L1 1.31 · L2 0.93 · L3 0.72 · check 0.24.noepoch/stark-s3at3992a56f9, gated by FAST 399; A/B on one binary each):f6445c8b6(FAST 452, gated: −2.48 s againstNOEPOCH_TREE_AHEAD=0, P8 10/10). The leaf programs are emitted beside the base. Leaf and node artifacts are built on the card during level 0, and each node program is emitted as soon as its children's artifacts exist. Base +0.01 s; every child proof verified.c9c29cb5b,NOEPOCH_BUILDER_FIRST, off): recursion −0.60 s in its A/B, but not by the mechanism pre-registered. It cannot pre-empt a leaf'smulti_prove, so the level-1 artifacts are still late; what it did was remove level 0's slow mode (≈ −0.4 s pooled over 451–453). Not the default.7da2a45d0; the default since08ecc4310, gated by FAST 456 and 457): recursion −0.29 s, level 0 −0.40 s. Level 0 was bimodal (5.2 or 5.8 s). In the slow runs, one leaf'smulti_proveheld the card for 1.2 s at about a third busy, starting at the instant the builder began emitting the level-1 programs on the global rayon pool (FAST 454's card trace). With a 4-thread host-only pool of its own, the stalled hold is gone in every run.NOEPOCH_EMIT_POOL=0restores the global pool.e029731f7; FAST 456): base −0.57 s (P8, 5 against 5 on one binary, no overlap), replicating WHIR's FAST 417 (−0.52 s).LAMBDA_VM_BUILDER_CONCAT=serialrestores the old copy.NOEPOCH_TREE_EAGER): recursion −0.06 s, because level 1 is limited by the card, not by waiting. The card trace (FAST 454) puts the idle inside the recursion's card holds at 1.2–1.8 s, 13–19 %.f6445c8b6(FAST 452); 35.03 s withNOEPOCH_TREE_AHEAD=0. Fan-in 4 was measured on the inline path (33.82–33.95 s there).a1a3a2c22through production's block verifier, all verified (bases 24.52–25.17 s), so the race fix (a) is shown on the plan's code.ElfConstants) and the per-(AIR, length) shapes can be reused across blocks.Memory: toward typical blocks (head
19fe5e9e5)A median mainnet block (25475471) is 9.78× this one, and the host peak grows about linearly with the block. The head carries three changes for that; none changes a proof byte.
drop_streamed_ops, default on;LAMBDA_VM_BLOCK_DROP_OPS=0keeps them)BLOCK_RECOMMIT_TOP_LEVELS)LAMBDA_VM_FIXED_TRACE_HASH=1)LAMBDA_VM_FIXED_TRACE_HASH=1andLAMBDA_VM_DETERMINISTIC_GRIND, two processes produce the same block proof (digestac406fc8…); a random-key control differs. Byte-identity gates on the block use this.What it is
VmProof. Every table is cut into instances of 2^21 rows, and KECCAK_RND into 2^16-row instances. There is no L2G table and no global proof; memory uses the monolithic PAGE argument.ResidencyMode::RecomputeLdeDevice: after Round 1 only each instance's root is kept. Each instance's fused task commits its trace on the device again, and the prover refuses the proof unless the new root equals the absorbed one, so the proof bytes are unchanged.RetainandRecomputeLdebehave as before.prover::block::prove_block/verify_block. The block verifier is the only one that accepts a chunked KECCAK_RND (AcceleratorShape::KeccakRndChunked). Every other verifier keeps KECCAK_RND to one table.LAMBDA_VM_GATE_PACKING(VRAM-gate packing admission),LAMBDA_VM_TABLE_TIMELINE(per-table timeline),LAMBDA_VM_RECOMMIT_TOP_LEVELS(kept top levels).WindowedTraceBuilder, and every full chunk of CPU, MEMW_R, MEMW_A, MEMW, LOAD, LT, SHIFT and STORE is committed as soon as it exists. Host tests show it equals the serial build.0fe8e4f2f; k = 3 before) areprove_block's default. The ELF data pages' preprocessed roots are computed on the device during execution.Base, block 25368371 on FAST (A/B on one binary; the epoch base is #1009's
prove_continuationon the same binary)e7406d6b5, FAST 396)At
e7406d6b5the block base is about 4.5 s below the epoch base (21.55 against 26.05). The first shared builder cut Round 1's span from 6.3 to 2.4 s, but its window builds sat on the executor's path (execute 1.35 → 6.9 s). With the rebuilt builder and the walk on its own thread, execute is 3.35 s and phase A 7.13 s, and the prove is 14.4 s against the serial build's 18.4 s. Nsight on FAST shows the fused phase is 98.6 % card-busy, so the block is card-bound. Most of that card time was the second hash, which kept top levels remove: phase B recomputes the LDE alone, and the openings rebuild each queried 8-leaf subtree and check it against the kept node. The architecture's main gain is in the recursion (above).S0 census: main 3.301 G elements, aux 1.036 G.
Format change: ECDAS chunked; KECCAK_RND and ECDAS heights capped
Landed at
23b3c8173(FAST 531 gates GREEN; FAST 533 A/B on 25512221: base −0.17 s, no effect on time, as expected; ECDAS cells −25 %, host peak −0.40 GiB).ECDAS is one row per double/add step (≈ 382 per ECSM call, ≈ 4.2 calls per transaction), so a median mainnet block
makes ≈ 420 k rows: one table at 2^19 proves 127.91 bits at DEEP batching on a ≈ 22.5 GiB device set, and a p90 block's
2^20 (126.91 bits, 44.7 GiB) no longer fits a 32 GiB card. The block now cuts ECDAS into instances of at most 2^17 rows;
a scalar multiplication may straddle two instances, its steps chaining only through the Ecdas bus, keyed by the call's
timestamp and the step's
(round, op).AcceleratorShape::KeccakRndChunkedis renamedBlockChunkedand lifts theone-table bound for ECDAS as for KECCAK_RND (count bounded by the sub-proof cross-check). New verifier constants, checked
in
verify_blockand in the tree'scheck_shape: every KECCAK_RND instance ≤ 2^16 rows (129.43 bits) and every ECDASinstance ≤ 2^17 (129.91 bits). This closes G1 (REV-JUDGE item 13, R-NOEPOCH-S3 F1-M1) for KECCAK_RND and ECDAS
only; every other table's height is still bounded by two-adicity alone, and the rest of G1's per-type list (ECSM,
KECCAK, the CPU family, the fixed tables) stays parked.
LAMBDA_VM_BLOCK_KECCAK_RND_LOG2takes 5..=16 (nooff).A block whose ECDAS fits one 2^17 table (25368371: 2^16) builds the same ECDAS table as before (one chunk of every step is the table
generate_optionalbuilt; by construction, not byte-compared on a block); 25512221 (2^18 today, 128.91 bits, 0.04under the minimum of record) now proves 2^17 + 2^16. The epoch and recursion verifiers (
Single) are unchanged.Format change: KECCAK and ECSM chunked; tree partition rule v2 (
5e4176961)Landed with the any-block target. Approved by the lead; Mauro to confirm. It is listed for the cryptography review as S-6 and S-9.
The caps. KECCAK is cut into instances of at most 2^18 rows (129.213 bits) and ECSM into instances of at most 2^17 (129.488 bits), as ECDAS is.
LAMBDA_VM_BLOCK_KECCAK_LOG2(2..=18) andLAMBDA_VM_BLOCK_ECSM_LOG2(2..=17) lower the caps, for tests.Partition rule v2 (
PARTITION_COST_MODEL = 2). Rule v1 pinned every chunk of a table to one leaf. Now the first instance keeps its seeded leaf and later chunks of KECCAK, ECSM and ECDAS fill by load. Without this, a p99 block's ≈ 15 ECDAS chunks overflowed the leaf cap and the plan was refused. The verifier derives the partition by the rule from the shape alone.Evidence:
Known cost (deferred): a table just over a power of two is cut after padding, so for example 2^20 + 1 KECCAK calls make 8 instances, about half of them padding. This matters only on keccak-heavy blocks.
LogUp: four interactions per aux column (opt-in,
LAMBDA_VM_ZF_LOGUP=k4)Default
pair= today's bytes.LAMBDA_VM_ZF_LOGUP=k4lets each base table commit four bus interactions per LogUp aux column instead of two, where that commits fewer extension columns: groups of four have degree 5, which blowup 4 admits, at the price of four composition parts instead of two. The rule is per table and verifier-side (⌈N/k⌉ aux columns + parts, ties keep pairs): KECCAK_RND 516 + 2 → 258 + 4, ECSM 290 → 145, ECDAS 194 → 97, CPU 10 + 2 → 5 + 4, MEMW_A 10 → 5; LT, STORE, MEMW_R, LOAD, PAGE and the small tables keep pairs. The LFM chips keep pairs; #1014 is untouched.ProofFormat.logup(stark), the group/accumulator emitters for k ≥ 3 (one body for the prover folder, the verifier folder and the IR capture), the host and device aux builds grouped by arity, the four-part composition split on the card (radix-2 twice) and its host mirror, compiled kernels for the k4 table programs, the knob atblock_base_options, and one new verifier refusal (parts > blowup).=pairthe 1× digest is ac406fc8 and every program id, kernel key and golden is unchanged (the small-tree id pin and the 44-program golden pass on the landing merge).Recursion: sibling proofs share the card through one VRAM gate (on by default)
Before: the recursion's card permit was a mutex, so one proof at a time ran inside
multi_prove. The card then sat idle while the holder ran its host stages: uploads, absorbs, queries. At the median block that was 24.6 s of card idle inside the holds (BIG 469).Now: every
multi_provein the tree admits its tables through one process-wide byte gate, and the artifact commit takes its bytes from the same gate. Sibling proofs overlap wherever their bytes fit. Three parts keep the gate's account matching the card:Retainprove keeps each table's main LDE, trace snapshot and tree on the card from its Round-1 commit until its fused task ends. Those bytes stay in the gate the whole time: the Round-1 task carries them past its own permit, and the fused task takes them over and is admitted only for the rest of its set.multi_proveat once.LAMBDA_VM_SHARED_VRAM_GATE=0restores the exclusive permit.LAMBDA_VM_SHARED_GATE_TRACE=1prints the gate's account (SGATElines). The gate acts only while the tree arms it for concurrent proofs, so the base and every single-proof path are unchanged.Bytes: unchanged, since the gate only reorders admission. The digest is
ac406fc8…, and every A/B arm has the same top program id.Measured:
a93018974against564dd02bb, 4 + 4):Tight claims (head
00871b913, on by default). A claim is a held part R plus a headroom H.LAMBDA_VM_SHARED_GATE_CLAIMS=wholerestores the first form (residents plus the largest table's whole set).278e6a8c6); no regression.Readout note: FAST 871's "claims in force" row read OUT because it counted from the optional trace, which that gate ran without. From each claim's own log line, every gate-on run peaked at 3 claims and the off runs had none.
Tests:
Disk spill: auto by default
What it does. Once Round 1 has committed an instance, its packed main trace can go to a spill file that phase B reads back ahead of its walks. The words that come back are the words that went out, so no proof byte depends on the policy.
LAMBDA_VM_BLOCK_SPILL=auto(default) |off|always|<GiB>(a resident budget for committed packed traces).How
autodecides (prover/src/block.rs,spill_target_bytes/spill_decision). A committed instance is spilled when the host's bytes, plus the reserve, plus the instance's own bytes would pass the target.memory.current, v1memory.usage_in_bytes) less its inactive file pages frommemory.stat, which the kernel reclaims first. Sincecf253d235; the charge alone counted the page cache, and at 1× on FAST it read 25.57 GiB against a working set of 17.09.LAMBDA_VM_BLOCK_SPILL_TARGET_GIB, if set;MemTotal, less 10 GiB. The limit is v2memory.max, or v1memory.limit_in_bytes, read at the process's cgroup path and then at the hierarchy root, which is what a container without a cgroup namespace sees. Since278e6a8c6: FAST 509 read 47.5 GiB on FAST's v1 limit, where the v2-only rule read 49.9;fb0dc0057, underautothey arm only once a spill is plausible (the host's bytes plus the reserve reach 85 % of the target) and then stay armed;alwaysand a fixed budget arm them from the start. At the median they never arm (waits 0.01 / 0.18 s, BIG 122). At the p75 they armed at 59.7 s, for a phase-A cost of +1.12 s against off; armed from the start, the same block cost +3.6 s (BIG 118).Measured, median block 25475471 on BIG (3 runs per arm, one binary per job):
auto(default)auto, 70 GiB targetalwaysalwaysalways(first measurement)ac406fc8…holds under the default (BIG 117) and underalways(BIG 114–116).Blocks up to full gas on BIG (128 GiB host, 120.69 GiB cgroup; one run each):
edddc6873035aef5d6035aef5d6autopurged at phase-a (121.19 → 82.32 GiB) and at the base (116.53 → 48.45). Its verifier passed cold and warm (66.59 / 61.61 s, pool high-water 17.39 GiB; BIG 125). With both purges forced on the pre-gate train it read 549.32 s at 111.30 GiB (BIG 1215).The allocator purge. The harness's jemalloc never purges (
dirty_decay_ms:-1), and a phase rarely reuses the pages the phase before it freed, so the host ratchets from phase to phase. Onearena.<all>.purgehands every arena's dirty pages back.autopurges only once the spill's budgets arm, so a block that fits printsALLOC PURGE <point>: skipped (auto, no memory pressure)and pays nothing (FAST 515 at 1×, BIG 123 at the median).What the cost follows. While the writer runs, the generators slow by ≈ 16 %. The writer's two passes over every spilled byte (digest, then the aligned copy) and the frees of the spilled buffers both contribute; BIG 115 and 116 could not split them further. At the median, phase A is bound by the generators during the walk, so the walk waits on them.
alwayshands ≈ 39 GiB to the writer during the walk, for +6–7 s.always: digest 22, aligned copy 28, pwrite 14. The pwrite runs at the disk's own rate (3.7 GiB/s raw). The BLOCK SPILL line prints the split.Measured and not landed (each has a FAILED-LEVERS row):
Tests:
always, a zero budget andautobuilds the resident traces;autoreads the host's working set, and the target reads v2 and v1 cgroup limits;autopurges only under memory pressure.